Papers with multimodal baselines
SpatialMath: Spatial Comprehension-Infused Symbolic Reasoning for Mathematical Problem-Solving (2026.findings-eacl)
Copied to clipboard
| Challenge: | Current models struggle to accurately decompose intricate visual inputs and connect perception with structured reasoning, leading to suboptimal performance. |
| Approach: | They propose a Spatial Comprehension-Infused Symbolic Reasoning Framework to integrate spatial representations into structured symbolic reasoning chains. |
| Outcome: | The proposed framework outperforms existing models in vision-intensive mathematical problems. |
MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations (P19-1)
Copied to clipboard
| Challenge: | Emotion recognition in conversations has gained popularity due to its potential applications. Until now, a large multimodal multi-party emotional conversational database containing more than two speakers per dialogue was missing. |
| Approach: | They propose to extend and enhance EmotionLines by combining 13,000 utterances from Friends dialogues with emotion and sentiment labels. |
| Outcome: | The proposed dataset contains about 13,000 utterances from 1,433 dialogues from the TV-series Friends. |
Open Your Model’s Eyes: Video and Context-Aware Multimodal Backchannel Prediction (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods for predicting backchannels rely on audio and text . existing methods omit visual cues and conversational contexts for accurate prediction . |
| Approach: | They propose a framework that leverages visual cues and conversational contexts to enhance backchannel prediction. |
| Outcome: | The proposed framework outperforms existing methods and simple multimodal baselines in recognizing complex backchannels such as empathy. |
MATCHED: Multimodal Authorship-Attribution To Combat Human Trafficking in Escort-Advertisement Data (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for human trafficking detection ignore the multimodal nature of online ads . sex trafficking is a pervasive crime exploiting individuals of all ages and genders . |
| Approach: | They propose to use multimodal authorship attributes to identify suspicious ads that combine text and images to improve vendor identification and verification tasks. |
| Outcome: | The proposed model outperforms existing methods for vendor identification and verification tasks using text-only, vision-only and multimodal training objectives. |
CI-AVSR: A Cantonese Audio-Visual Speech Datasetfor In-car Command Recognition (2022.lrec-1)
Copied to clipboard
Wenliang Dai, Samuel Cahyawijaya, Tiezheng Yu, Elham J. Barezi, Peng Xu, Cheuk Tung Yiu, Rita Frieske, Holy Lovenia, Genta Winata, Qifeng Chen, Xiaojuan Ma, Bertram Shi, Pascale Fung
| Challenge: | In-car smart assistants should be able to process general as well as car-related commands and perform corresponding actions, which eases driving and improves safety. |
| Approach: | They propose a dataset for in-car command recognition in the cantonese language with both video and audio data. |
| Outcome: | The proposed model can achieve a considerable quality on the clean test set, but the speech recognition quality on noisy data is still inferior. |
Not all Fake News is Written: A Dataset and Analysis of Misleading Video Headlines (2023.emnlp-main)
Copied to clipboard
| Challenge: | Social media platforms are used by half of U.S. adults for everyday news consumption. |
| Approach: | They propose to analyze video headlines and whether annotators believe the headline is representative of the video’s contents. |
| Outcome: | The proposed dataset analyzes video headlines and explains why annotators view a video as misleading. |
MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation (2025.findings-emnlp)
Copied to clipboard
| Challenge: | MEXA is a training-free framework that performs modality- and task-aware aggregation of multiple expert models to enable effective multimodal reasoning across diverse domains. |
| Approach: | MEXA is a training-free framework that performs modality- and task-aware aggregation of multiple expert models. |
| Outcome: | MEXA performs modality- and task-aware aggregation of multiple expert models . it generates interpretable textual reasoning outputs and reasons over them using a Large Reasoning Model (LRM) MEX A consistently delivers performance improvements over strong multimodal benchmarks . |
Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing MLLMs are strong at understanding single plots, but struggle with multi-step reasoning . Existing approaches to manage context in chart reasoning include text-based chain-of-thought prompting . |
| Approach: | They propose a hierarchical visual agent framework that iteratively constructs a working context in an image–text space. |
| Outcome: | The proposed framework improves on strong multimodal baselines. |